Papers by Nisansa de Silva
Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial Regularization (2020.coling-main)
Copied to clipboard
Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li
| Challenge: | Existing graph autoencoders and its variants have been used for node embedding . a new method is proposed to model consistency across different views of networks . |
| Approach: | They propose a network embedding method which enforces latent representations to be consistent across different views of networks by incorporating a multiview adversarial regularization module. |
| Outcome: | The proposed method compares favorably with the state-of-the-art methods on benchmark datasets and on a real-world application. |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022.tacl-1)
Copied to clipboard
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofetoluwa Adeyemi
| Challenge: | Lower-resource corpora have systematic issues, including mislabeled or nonstandard/ambiguous language codes. |
| Approach: | They manually audit the quality of 205 language-specific corpora released with five major public datasets. |
| Outcome: | The results show that lower-resource corpora have systematic issues even for non-proficient speakers. |
Semantic Oppositeness Assisted Deep Contextual Modeling for Automatic Rumor Detection in Social Networks (2021.eacl-main)
Copied to clipboard
| Challenge: | Social networks face a major challenge in the form of rumors and fake news . rumor detection is suboptimal due to its rapidity and spread of information . |
| Approach: | They propose a semantic oppositeness model that captures elements of discord . they show that it is more resistant to variances introduced by randomness . |
| Outcome: | The proposed model achieves state-of-the-art on rumor detection task with extensive experiments on recent data sets. |
Some Languages are More Equal than Others: Probing Deeper into the Linguistic Disparity in the NLP World (2022.aacl-main)
Copied to clipboard
| Challenge: | Linguistic disparity in the NLP world is widely acknowledged, but the reasons behind it are rarely discussed within the field. |
| Approach: | They propose to categorise languages based on speaker population and vitality . they also analyse the distribution of language data resources and amount of NLP/CL research . |
| Outcome: | The proposed model identifies the reasons for the disparity and suggests ways to overcome it. |
Improving the Quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics (2025.emnlp-main)
Copied to clipboard
| Challenge: | Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from web-mined corpora. |
| Approach: | They propose to rank parallel sentences using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) . previous research has shown that the choice of multiPLM significantly impacts the quality of the filtered parallel corpus. |
| Outcome: | The proposed methods reduce disparities between multiPLMs while producing better results. |